Papers with multi-document summarization task

5 papers
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation (2024.emnlp-main)

Copied to clipboard

Challenge: Various attempts to correct noisy data in the construction process have been made, but human annotation is expensive and time-consuming.
Approach: They propose to use large language models for data annotation to imitate human annotation and classify unrelated documents from a multi-document summarization task.
Outcome: The proposed method imitates human annotation and classifies unrelated documents from the Multi-News dataset.
Unsupervised Aspect-Based Multi-Document Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing methods for opinion summarization are expensive and do not deal with contradictory statements.
Approach: They propose an unsupervised abstractive summarization neural system that generates short summaries of reviews in a vector space.
Outcome: The proposed system can generate short summaries of user-generated reviews in a short paragraph, while nobody reads all reviews.
What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to summarize clinical narratives are lacking.
Approach: They propose to generate a paragraph that tells the story of a patient's hospitalization . they analyze a dataset of 109,000 hospitalizations and their corresponding summary proxy .
Outcome: The proposed model is based on a dataset of 109,000 hospitalizations and their corresponding summary proxy.
Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles (2020.emnlp-main)

Copied to clipboard

Challenge: Multi-XScience is a dataset construction protocol that favours abstractive modeling approaches.
Approach: They propose a large-scale multi-document summarization dataset that is based on articles and lexical databases and WordNet synonymy information to generate related-work sections of a paper.
Outcome: The proposed method is based on lexical databases and WordNet synonymy information to write related work sections of a paper based upon their abstract and the articles they reference.
Benchmarking LLMs on Semantic Overlap Summarization (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are the most capable text generation models in a variety of tasks and fields.
Approach: They benchmark Large Language Models (LLMs) on SOS and introduce PrivacyPolicyPairs (3P) a dataset of 135 high-quality privacy policy documents is used to evaluate the model.
Outcome: The proposed dataset complements existing resources and broadens domain coverage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations